Back

Journal of Genetics and Genomics

Elsevier BV

Preprints posted in the last 90 days, ranked by how well they match Journal of Genetics and Genomics's content profile, based on 38 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
CellClick: an interactive platform for adjustable and accurate cell type annotation in single-cell and spatial omics data

Shi, L.; Dai, M.; Zhang, Y.-b.; Wu, S.; Wang, M.; Wang, X.-j.

2026-06-03 bioinformatics 10.64898/2026.06.01.727775 medRxiv
Top 0.1%
4.6%
Show abstract

Single-cell omics and spatial omics technologies are nowadays widely used in biological and medical research. In both single-cell and spatial omics data analysis, accurate cell type annotation is a key step for downstream analysis and scientific discoveries. However, high-quality cell annotation usually requires multiple rounds of manual analysis for result refinement, which poses great challenges to most researchers. Here, we present CellClick, an interactive platform for convenient and accurate cell type annotation in single-cell and spatial omics data. CellClick provides Data Preprocessing, Data Visualization, Cell Annotation, Annotation Validation, and Cell Reannotation modules, which facilitate automatic or user-guided cell selection and annotation. The feasibility of using CellClick to generate more accurate cell annotation results was exemplified by both scRNA-seq and spatial transcriptomics data.

2
A Korean pangenome reference of 14 healthy individuals supports structural variant analysis in disease genomes

Shin, D.-H.; Jeon, J.; Joe, S.; Jeon, Y.; Yang, J. O.; Bhak, J.; Baek, S. A.; Byun, G.; Shin, E.-S.; Kwon, Y.; Choi, H.-J.; Kim, J.-H.; Haam, K.; Yoo, J.; Song, K. J.; Mok, J.; Jeon, S.; Jeong, H.; Bhak, J.

2026-07-09 genetic and genomic medicine 10.64898/2026.07.06.26357367 medRxiv
Top 0.1%
2.7%
Show abstract

Here, we present the first graph-based Korean Pangenome Reference (K-PanRef), constructed from 14 healthy Korean individuals. K-PanRef comprises 13 high-quality diploid Korean genome assemblies (mean QV ~62.0) and KOREF1-G-TTAGGA, the first complete Korean reference genome. Integration of these assemblies generated a ~3.2-Gb pangenome graph containing ~39.3 million nodes and ~53.8 million edges, with the accumulation of common sequences (frequency [≥]10%) reaching a plateau. Additionally, K-PanRef contains ~4.3 million Korean-specific small variants and ~76.0 thousand Korean-specific SVs absent from the Chinese and human pangenome references, improving the representation of Korean genetic diversity relative to these references. To evaluate its utility for short-read-based SV analysis, we genotyped 75 whole-genome sequencing (WGS) samples, including 15 patients with early-onset myocardial infarction (MI). Although constructed entirely from healthy genomes, K-PanRef supported the identification of putative disease-relevant SVs in this exploratory application. K-PanRef-based genotyping identified ~95.6 thousand small variants and 820 SVs observed only in the early-onset MI samples. Among the early-onset MI-group SVs, 491 were absent from public databases, suggesting that they may represent previously unrecognized candidate variants related to early-onset MI. Of these, 164 SVs overlapped 134 genes, of which 89 had reported associations with 42 cardiovascular diseases or traits, including eight genes previously linked to MI. Together, these results establish K-PanRef as a valuable resource for representing Korean genetic diversity and enabling more comprehensive discovery of population-specific and novel putative disease-relevant variants from short-read sequencing data.

3
CN-RNN: a Deep Learning Framework for Copy Number Variation Detection with Exome Sequencing Data

Wang, D.; Qin, F.; Bao, W.; Bacher, R.; Chung, D.; Lu, Q.; Efron, P. A.; Cai, G.; Xiao, F.

2026-05-15 genetics 10.64898/2026.05.13.724920 medRxiv
Top 0.1%
2.7%
Show abstract

Copy number variations (CNVs) are major structural genomic variants that contribute to a wide range of human diseases. Accurate detection of CNVs from whole-exome sequencing (WES) data has been a long-sought goal for clinical and population genetic studies. Despite recent progress, existing WES-based CNV callers still suffer from high false-positive rates and reduced recall for short-length variants, and current deep learning methods have not fully used complementary information in region-level genomic features. Here we present CN-RNN, a deep learning-based CNV caller for WES data. The model combines a bidirectional long short-term memory (BiLSTM) branch that captures local depth changes and contextual dependencies across neighboring exons with a parallel multi-layer perceptron (MLP) branch that encodes region-level metadata such as GC content, mappability, and exon length. CN-RNN was trained on the Autism Sequencing Consortium (ASC) parent-child trio cohort using the Mendelian rule of inheritance to ensure high-quality training sets. It was evaluated across three independent datasets, in which we showed that CN-RNN outperformed existing WES-based CNV callers and deep learning methods. CN-RNN offers a scalable, accurate tool for CNV profiling in WES-based studies and supports broader application of CNV analysis in population and clinical research. CN-RNN is available at https://github.com/FeifeiXiao-lab/CN-RNN.

4
ExMODE: A Multi-Omics Repository for Extremophile Adaptation and Bioprospecting

Li, D.; Ma, K.; Zhang, Y.; Wang, J.; Cui, Z.; Li, X.; Wang, W.; Tong, J.; Guo, Y.; Wang, Z.; Zeng, P.; Wang, J.; Xu, X.; Zhang, N.; Zhang, Y.; Chen, J.; Hu, Q.; Yang, W.; Li, Z.; Yang, T.; Du, W.; Xu, Z.; Yue, Z.; Wang, J.; Fan, G.; Zhang, W.; Xu, X.; Huo, L.; Wei, X.; Meng, L.; Liu, S.

2026-04-29 microbiology 10.64898/2026.04.27.720953 medRxiv
Top 0.2%
2.4%
Show abstract

Extreme environments, though hostile to most life forms, host specialized extremophile communities that have redefined biological cognition and emerged as vital biotechnological resources, with their unique adaptive traits and bioactive molecules driving advances in multiple scientific and industrial fields. However, research on extremophiles is hindered by limitations in culture-based methods, fragmented multi-omics data with non-uniform annotation standards across repositories, the lack of cross-extreme comparative research in existing resources, and the singularity of data dimensionality that neglects key structural information, all of which restrict the functional interpretation of extremophile microbes and the exploitation of their bioprospecting potential. To tackle these challenges, we developed ExMODE (https://db.genomics.cn/exmode/), a comprehensive multi-omics database platform dedicated to extremophiles. It centrally integrates multi-omics data from diverse extreme habitats with a standardized annotation framework, resolving data fragmentation and enabling systematic cross-environment comparative analyses to elucidate extremophile adaptive mechanisms. Moreover, ExMODE aggregates multi-dimensional datasets including genes, genomes, secondary metabolite sequences and protein structures, overcoming the constraints of single-dimensional data and significantly improving the efficiency of biotechnological resource discovery from extreme microorganisms.

5
Long-read sequencing reveals transposable element-derived chimeric transcripts at zygotic genome activation in mammalian embryos

Kawakami, S.; Kitao, K.; Ikeda, S.; Honda, S.

2026-05-28 developmental biology 10.64898/2026.05.25.727629 medRxiv
Top 0.2%
2.3%
Show abstract

BackgroundTransposable elements (TEs) are mobile genomic sequences that constitute one-third to one-half of the mammalian genome. Recently, TEs have been recognized for their important roles as cis-regulatory elements. TEs are broadly activated during zygotic genome activation (ZGA) in mammalian embryos, where they function as alternative promoters of host genes and drive the transcription of chimeric transcripts. However, the construction of comprehensive chimeric transcript databases based on short-read sequencing remains limited due to the repetitive and abundant nature of TEs in the genome. Here, we used long-read RNA sequencing to construct a comprehensive dataset of chimeric transcripts expressed in ZGA mouse and bovine embryos. ResultsWe identified 11,996 and 4,755 chimeric transcripts variants derived from 2,695 and 1,200 host genes in mouse and bovine, respectively, exceeding the numbers reported in previous short-read-based studies. Among them, 114 orthologous pairs produced chimeric transcripts in both species. Gene Ontology analysis revealed significant enrichment of terms related to transcriptional regulation and protein modification in mouse, whereas no terms were significantly enriched in bovine. Assessment of the protein-coding potential of the TE-driven transcripts using predicted open reading frames (ORFs) revealed that the proportion of "Protein-coding" transcripts was lower, whereas that of "LncRNA" (long non-coding RNA) was higher compared with all transcripts in both species. Among the ORFs classified as "Protein-coding", comparison with canonical ORFs revealed a tendency for the N terminus to be truncated while the C terminus remained intact in both species. TE-derived promoters used in mouse were enriched for mouse-specific TEs, whereas those in bovine were enriched for older TEs conserved among eutherians. In addition, long-read sequencing detected a greater number and proportion of TEs used as promoters in mouse and bovine than short-read sequencing. Although motif analysis identified KLF5 and OTX2 binding sites upstream of TE-derived promoters in both species, the specific TEs containing these motifs differed between the two species. ConclusionsThis study presents the first long-read sequencing analysis of chimeric transcripts in mammalian embryos in two species. Our approach revealed the functional similarities of chimeric transcripts between species, as well as species-specific differences in their TE compositions.

6
OneGenomeRice (OGR): A Genomic Foundation Model for Rice

Qian, B.; Liang, C.; Qin, C.; Liu, C.; Zhang, C.; Xu, C.; Li, D.; Xue, G.; He, H.; Zhang, H.; He, H.; Chen, D.; Xu, J.; Zhang, J.; Sun, J.; Shang, L.; Jiang, J.; Xia, K.-k.; Zhong, L.; Chen, L.-l.; Fan, L.; Liu, L.; Qin, M.-m.; Li, Q.; Zhu, S.; Ma, S.; Liu, S.; Zhang, S.; Fu, S.; Wei, T.; Xu, X.; Jia, X.; Xu, X.; Jing, Y.; Xu, Y.; Zhao, Y.; Xue, Y.; Guo, Y.; Xiao, Z.; Li, Z.; Li, Z.; Yue, Z.; Deng, Z.

2026-04-23 genomics 10.64898/2026.04.21.719822 medRxiv
Top 0.3%
1.7%
Show abstract

The transition of genomics to a predictive intelligence discipline is driven by the advent of genomic foundation models. While substantial progress has been observed in human-centric models, plant genomics, particularly for the staple crops, remains hindered by a lack of models. Here we introduce OneGenomeRice (OGR), a genomic foundation model for rice (Oryza sativa) engineered by a Mixture of Experts (MoE) transformer architecture with 1.25-billion-parameters. OGR was pre-trained on a genomic dataset comprising 422 high-quality genomes of cultivated and wild rice. A comprehensive benchmark, including short-sequence motif identification, long-range regulatory modeling, single-nucleotide resolution prediction, selective sweep detection and subspecies classification, demonstrated that OGR significantly outperforms existing state-of-the-art plant or all-life genome models in 11 categories. The model was also further used for several downstream applications, such as introgression between indica and japonica subspecies using embedding-based supervised classification, agronomy trait-associated functional loci through attention-derived importance signals, and gene expression prediction of DNA sequences etc. These results indicate OGR being a promising foundational computational infrastructure for functional genomics and precision breeding of rice.

7
A Dual-Locus-Targeting Strategy to Enhance CRISPR/Cas9-mediated CFTR Replacement via Helper-Dependent Adenoviral vector in porcine genome

Chen, Z. R.; Zhou, Z. P.; Duan, R. C.; Wong, A.; Grasemann, H.; Bear, C.; Hu, J.

2026-06-11 genetics 10.64898/2026.06.10.731381 medRxiv
Top 0.3%
1.6%
Show abstract

Gene therapy has been the subject of extensive research following the advent of gene-editing technologies. Genetic disorders with difficult-to-target tissues, such as cystic fibrosis (CF), still face many challenges in developing efficacious gene therapy. The potential universal approach of gene replacement involves inserting a functional CFTR gene after generating DNA double strand breaks using gene editors such as CRISPR/Cas9. However, this strategy has not achieved clinical significance, as CRISPR/Cas9-mediated integration of CFTR is limited primarily by the infrequent activity of the homology-directed repair (HDR) pathway. To circumvent this limitation and improve CFTR transgene integration and expression, we explored a method of adding a second integration site, which we termed the dual-locus-targeting method. Using a helper-dependent adenoviral vector (HDAd)-delivered CRISPR/Cas9 system in porcine epithelial cells, we found that sequential delivery of two vectors, one targeting the CFTR locus and the other the genomic safe harbour site GGTA1, enhanced the integration efficiency of lacZ and CFTR donor genes to 16.5% and 3.4%, respectively. These results demonstrated a potential strategy to improve the efficacy of CFTR replacement for the development of a universal and permanent gene therapy treatment for CF lung disease. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=76 SRC="FIGDIR/small/731381v1_ufig1.gif" ALT="Figure 1"> View larger version (17K): org.highwire.dtl.DTLVardef@1774590org.highwire.dtl.DTLVardef@1782915org.highwire.dtl.DTLVardef@1d13b12org.highwire.dtl.DTLVardef@17d3f93_HPS_FORMAT_FIGEXP M_FIG C_FIG

8
A high-quality, chromosome-scale genome assembly of the shade-tolerant wild rice, Oryza granulata

Zhang, F.; Yang, Y.-h.; Li, W.; Shi, C.; Zhu, X.-g.; Gao, L.-z.

2026-05-01 bioinformatics 10.64898/2026.04.28.721348 medRxiv
Top 0.4%
1.5%
Show abstract

Oryza granulata Nees et Arn. ex Watt, a diploid wild rice (GG genome), possesses exceptional shade tolerance and is a key genetic resource for rice improvement. However, previous genome assemblies lacked continuity and completeness. Here we present a chromosome-scale reference genome of O. granulata using PacBio SMRT (113x), Hi-C (95x), and Illumina sequencing. The final assembly is ~764.24 Mb, with a scaffold N50 of ~59.32 Mb, and ~96.47% of the sequence anchored to 12 chromosomes. BUSCO completeness is ~98.6%. We annotated ~42,064 protein-coding genes, of which ~95.39% were functionally annotated, along with ~73.46% repetitive elements. The genome assembly and raw sequencing data are available at NGDC (PRJCA061980), NGDC GSA (CRA068332), and NGDC GWH (GWHISVE00000000.1). This high-quality genome will serve as a fundamental resource for evolutionary genomics, conservation biology, and breeding of shade-tolerant rice cultivars.

9
Translational bioinformatics and machine learning framework for biomarker discovery, disease prediction, and patient profiling for precision medicine

Ahmed, Z.; Govindareddy, P.; DeGroat, W.; Narayanan, R.; Peker, E.; Zeeshan, S.

2026-05-27 genetic and genomic medicine 10.64898/2026.05.23.26353961 medRxiv
Top 0.4%
1.2%
Show abstract

Precision medicine aims to advance our ability from a "one-size-fits-all" approach to personalized and predictive healthcare across diverse populations. It promotes integration of multi-omics and phenotypic data to understand disease mechanisms and discover novel biomarkers and risk factors, which could be used to predict and prevent critical diseases in individual patients across diverse populations. The potential implications of precision medicine approach can accelerate our ability to classify patients at higher risk of developing critical diseases, improve diagnostic capabilities, develop deeper understanding of individual risk, investigate racial differences and demographic characteristics, and find relationships between genetic variants, expressions, and diseases. This study focuses on implementing an innovative and data driven framework of translational bioinformatics and Machine Learning (ML) techniques to analyze multi-omics, including RNA-seq and Whole-Genome Sequencing (WGS) data, generated using blood samples of randomly consented patients. First, we utilized bioinformatics pipelines to identify differentially expressed genes and their pathogenic and likely pathogenic variants for the downstream data analysis, annotation, and visualization. Then, applied a nexus of ML models for multi-omics biomarker discovery, disease prediction, density-based clustering, single-patient profiling, and pathogenicity classification. WGS data analysis supported the exploration of genetic variation and diversity among patients to identify known and novel biomarkers, whereas RNA-seq data analysis improved our understanding of functional and biological pathways that underlying disease states. We classified and clustered pathogenic variants and expressions across various genes and discovered numerous diseases leading risk factors. Our results include gene-disease associations and captured common pathways across the broader population, demonstrating a level of sensitivity and accuracy that has broad clinical implications. We validated our results through clinical records, and state of the science literature. This study delves into the strengths of multi-omics data integration and capabilities of ML application in genetically diverse and complex patient cohorts. Our approach has the potential to elucidate complex gene-disease interactions for genetically diverse populations, which can support earlier diagnoses for patients in many disease realms.

10
Integrative genomics reveals shared and stress-specific adaptive pathways underlying acidic soil-associated metal toxicity in rice

Jaiswal, S.; Kumari, A.; Singh, B. K.; Kumar, K.; Kumar, S.; Kumar, S.; Kaur, S.; Prakash, N. R.; Baiswar, P.; Bharati, A.; Talukdar, M.; Behera, S.

2026-06-03 genomics 10.64898/2026.05.31.729167 medRxiv
Top 0.5%
1.1%
Show abstract

Soil acidity-associated toxicities of aluminum (Al), cadmium (Cd), and manganese (Mn) severely constrain rice productivity in upland ecosystems. To investigate the genomic basis of adaptation to acidic soil-related metal stress, we conducted an integrated meta-QTL (M-QTL) and functional genomics analysis in rice. Meta-analysis of 681 QTLs and MTAs from 53 QTL mapping and GWAS studies identified 79 robust M-QTLs, including ten overlapping regions associated with Al-, Cd-, and Mn-responsive traits. A multi-criteria prioritization framework identified 98 candidate genes supported by positional overlap, transcriptomic recurrence, and functional annotation, enriched for ion transport, detoxification, and redox regulation pathways. M-QTL10.9 emerged as a major hotspot enriched for glutathione-S-transferase genes, whereas M-QTL9.5 contained the highest density of prioritized candidates linked to Al and Cd responses. Comparative physiological & biochemical analyses of the contrasting rice genotypes Sahasarang and IR64 revealed genotype-dependent differences in antioxidant responses, metal partitioning, metabolic regulation, and cell wall remodeling under individual and combined metal stresses. Expression profiling of prioritized candidate genes, including OsACO family genes, OsZIP10, and OsGSTU10, further revealed genotype-dependent transcriptional divergence under combined stress. The identification of overlapping M-QTLs across Al, Cd, and Mn datasets suggests both shared and stress-specific adaptive responses to acidic soil-associated metal stress in rice.

11
In Vivo Spatial Transcriptomics for Bleeding-free Profiling Human Internal Organs

Sun, H.; Guo, F.; Zhao, X.; Wan, Y.; Zhang, X.; Sun, J.; He, X.; Gai, B.; Xiong, C.; Ma, Y.; Qu, J.; Li, P.; Gao, F.; Zhao, X.; Ji, X.; Yang, Z.; Mak, L.-Y.; Yap, Y. H.; Ke, J.; Shi, P.

2026-07-09 genetic and genomic medicine 10.64898/2026.07.06.26357355 medRxiv
Top 0.5%
1.1%
Show abstract

Despite the significant technical advancement in spatial transcriptomics, its clinical usage is largely untapped. Here, we develop an integrated system, ENDO-Genome, for minimally invasive in-body transcript sampling to facilitate live spatial transcriptomic analysis of human internal organs. This is achieved by integrating a nanoarrayed biochip with existing endoscope to perform pressure-sensor-calibrated "Touch & Go" RNA extraction directly from human internal organs, including the highly vascularized liver or kidney, without the need for tissue biopsy, voiding any bleeding risks. By a demonstration using gastrointestinal endoscopy, multiplexed landscape of 55 mRNA transcripts was obtained from multiple locations of human intestinal tract via a 5-minute operation in routine examinations. Benefiting from a sequencing-free approach, each assay costs less than 10 US dollars. For the clinical study involving 15 Crohn' s disease (CD) patients, no complication case was reported out of 47 ENDO-Genome operations, showcasing the gentle deposition and excellent safety of the technique. The live spatial transcriptomics provides direct in vivo pictures of the heterogenous spatial transcriptional programs underlying CD pathological response at different intestinal locations, revealing distinct ileal phenotypes. This is manifested by unique microscale scattering of inflammation gene clusters, along with the discovery of a tissue-specific cooperative mechanisms between inflammation and RNA methylation regulations at single- or multi-cell scales.

12
Multi-layered characterization of ~700,000 conserved noncoding elements within the human genome

Fibi-Smetana, S.; Fernandez-Mendoza, F.; Taher, L.

2026-06-17 evolutionary biology 10.64898/2026.06.16.732547 medRxiv
Top 0.5%
1.1%
Show abstract

Conserved noncoding elements (CNEs) have been extensively studied for their roles as regulatory elements, particularly enhancers. However, the advent of technologies like ChIP-seq and ATAC-seq has shifted research focus away from comparative genomics. Here, we leveraged data from large-scale projects like ENCODE to address the resulting gap in the comprehensive functional characterization of CNEs. We first derived a set of [~]700,000 CNEs in the human genome from a 470-way mammalian alignment. Phylogenetic inference identified [~]670,000 conserved elements within primates and [~]240,000 conserved elements across mammals. Our functional genomic analysis revealed that, irrespective of their level of conservation, approximately one third of CNEs exhibit concurrent chromatin accessibility and H3K27 acetylation in at least one of 19 examined tissues and cell lines and thus, are likely to act as cis-regulatory elements. Extrapolating these data to additional tissues and cell lines suggested that [~]40% of the CNE repertoire possesses cis-regulatory potential. Moreover, we found that the 3D organization of CNEs is non-random; specifically, CNEs are preferentially located toward the centers of topologically associating domains. CNE co-activation networks derived from chromatin accessibility and active histone marks revealed that evolutionary constraints acting on CNEs functioning as cis-regulatory elements reflect not only their isolated individual role, but their topological context. To summarize, we have generated a novel catalog of CNEs annotated with empirical cis-regulatory evidence. While evolutionary constraint and regulatory function are clearly linked, a comprehensive understanding of their interplay remains elusive. This resource provides a foundation for exploring this relationship systematically. Significance statementPrevious studies have investigated conserved noncoding element (CNE) evolution, epigenomic landscapes, and 3D genome organization separately, yet a systematic framework integrating these dimensions has been lacking. Here, we identified [~]700,000 CNEs, including [~]240,000 CNEs deeply conserved across mammals, and show that a large fraction display enhancer-associated epigenomic signatures and are preferentially enriched within TAD centers, highlighting their regulatory relevance. By generating and analyzing this CNE catalog within an integrated evolutionary, epigenomic, and 3D genomic context, our study bridges a critical gap and provides a comprehensive resource to better understand regulatory architecture and its potential contribution to disease-associated variation.

13
Integrating spatial and single-cell multi-omics analysis of induced pluripotent stem cell-derived cervical adenocarcinoma model

Kamata, S.; Taguchi, A.; Iuchi, H.; Ikeda, Y.; Maruyama, R.; Nakanishi, Y.; Sugi, T.; Okuma, Y.; Kobayashi, O.; Tomita, N.; Yoshimoto, D.; Wang, L.; Moritsugu, N.; Takahashi, C.; Tagami, M.; Matsunaga, H.; Okayama, T.; Manabe, R.-i.; Kiyotani, K.; Ikeo, K.; Okazaki, Y.; Kiyono, T.; Masuda, S.; Hamada, M.; Takeyama, H.; Kawana, K.

2026-05-06 cancer biology 10.64898/2026.05.01.722143 medRxiv
Top 0.5%
1.1%
Show abstract

Human papillomavirus 18 (HPV18) preferentially infects cervical stem cell-like cells and is strongly associated with adenocarcinoma. However, the mechanisms underlying differentiation into cervical adenocarcinoma remain unclear due to the lack of appropriate experimental models. We aimed to establish a model of HPV18-associated cervical adenocarcinoma and elucidate its molecular and cellular differentiation mechanisms. HPV18 E6/E7 were introduced into induced pluripotent stem cell-derived reserve cell-like cells (iRCs) to generate tumor models. Spatial transcriptomics and single-cell multi-omics analyses were performed to integrate histological and molecular data. A distinct component (Gland_A) exhibited morphological and immunohistochemical features of cervical adenocarcinoma and was efficiently induced in iRC-18 tumors. Gland_A showed increased chromatin accessibility and elevated expression of FOXA1, FOXA2, and ALDH1A1. Analysis of clinical samples confirmed enrichment of ALDH1A1 in HPV-associated adenocarcinomas. This model recapitulates key features of HPV18-associated cervical adenocarcinoma and provides insights into its differentiation mechanisms.

14
Whole-genome duplication underlies conserved sexually biased expression of meiotic cohesin genes unique to the teleost fish lineage

Niwa, T.;Kikuchi, M.;Tanaka, M.

2026-06-27 Developmental Biology 10.64898/2026.06.26.731870 medRxiv
Top 0.6%
1.0%
Show abstract

Meiosis is a fundamental process in producing both sperm and eggs, yet recombination landscapes often exhibit sexual differences, known as heterochiasmy. Since meiotic proteins are generally expressed in both sexes, the molecular mechanism driving heterochiasmy remains elusive. The -kleisin subunit gene of meiotic cohesin, Rec8, is expressed bisexually in mammals, while its putative teleost ortholog, rec8a, is expressed in a female-biased manner, presumably due to the presence of its paralog originating from the teleost-specific whole-genome duplication (TGD). Here, we elucidated the evolutionary history and expression dynamics of -kleisin genes across teleost lineages. Through comprehensive phylogenetic and synteny analyses, we revealed that major teleost lineages retain two copies of rec8 and rad21, with rec8 loci experiencing drastic chromosomal rearrangements immediately after the TGD. Using in situ hybridization and single-cell transcriptome data in medaka and zebrafish, we demonstrated a conserved sexually biased expression pattern: rec8a is predominantly female-biased, whereas rec8b exhibits male-biased expression during gametogenesis. Furthermore, comparative epigenetic analyses revealed that the conserved sexually biased expression is driven by lineage-specific cis-regulatory elements, rather than conserved ones. Motif analyses imply that regulatory rewiring by transcription factors, including foxl2l in particular, might have played a crucial role in the establishment and maintenance of this paralog divergence. Our findings highlight how whole-genome duplication and subsequent genomic and epigenetic rewiring subdivided the bisexual function of rec8, offering insights into sexually distinct meiotic regulation. HighlightsO_LITeleosts possess a unique -kleisin repertoire originating from the TGD. C_LIO_LITeleost rec8 paralogs exhibit conserved sex-biased expression during meiosis. C_LIO_LIDrastic genomic rearrangements after the duplication rewired the teleost rec8 loci. C_LIO_LIThe conserved expression pattern is governed by lineage-specific CREs. C_LIO_LIThose CREs harbor similar types of TFBSs such as Fox-family TFs. C_LI Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=94 SRC="FIGDIR/small/731870v1_ufig1.gif" ALT="Figure 1"> View larger version (22K): org.highwire.dtl.DTLVardef@c8c84dorg.highwire.dtl.DTLVardef@1d65668org.highwire.dtl.DTLVardef@c2d732org.highwire.dtl.DTLVardef@1be54a2_HPS_FORMAT_FIGEXP M_FIG C_FIG

15
A comprehensive DNA methylation atlas for the Chinese population through nanopore long-read sequencing of 106 individuals

Li, Y.; Jiang, T.; Qian, L.; Wang, Y.

2026-04-23 genomics 10.64898/2026.04.20.719515 medRxiv
Top 0.6%
1.0%
Show abstract

DNA methylation constitutes the primary epigenetic language mediating organismal phenotypic plasticity. Establishing a cohort-level genomic methylation landscape featuring wide geographical diversity is fundamental for dissecting its genetic and environmental attributes. Leveraging nanopore sequencings strength in genome-methylome co-sequencing, we generated a whole-genome, haplotype-resolved methylation atlas for 106 individuals from 19 provinces across China. The atlas identified 27,609,354 CpG sites genome-wide, with notably more informed gene proximal regions and CpG islands compared to whole-genome bisulfite sequencing. Detailed analyses revealed genomic structural variants as a pervasive covariate of DNA methylation, with a remarkable 2-fold compensation effect found in genome-wide heterozygous deletions. On the other hand, habitat altitude is found to be a strong environmental determinant of DNA methylation. We established a quantitative relationship between altitude and methylation states and identified a gene set strictly responsive to altitude differences, revealing epigenetically regulated genes such as PRDM16, EPHB2 and WNT7A. The methylation atlas provides a reference resource to facilitate further explorations into human epigenetics.

16
Killiverse: an interactive multi-omics web resource for killifish

Mittal, A.; Singh, P. P.

2026-06-21 genomics 10.64898/2026.06.16.731504 medRxiv
Top 0.6%
1.0%
Show abstract

BackgroundKillifish have emerged as valuable vertebrate model systems for investigating several disciplines including aging, regeneration, and developmental biology. Multi-omics datasets are increasingly being generated for killifish. However, their reuse remains limited due to computational challenges, largely due to the lack of accessible resources creating a bottleneck in widespread adoption of the killifish model. To address this, we developed Killiverse, a web resource for quick and intuitive exploration of multi-modal omics data dedicated to the model organism. ResultsKilliverse is an interactive, no-code, web-based platform designed for exploration of killifish multi-omics data. The platform aggregates a growing list of datasets including bulk transcriptomes, single-cell and single-nucleus transcriptomes, proteomes, and lipidomes processed through standardized pipelines and genome assemblies. Killiverse supports customized visualization and enables cross-study and cross-species analysis. It provides ortholog mapping to several established model organisms. By combining low-code software development with modern cloud technologies, the platform delivers a scalable browser-accessible application for the community. ConclusionsKilliverse enables rapid hypothesis development through the identification of patterns across studies and species. The ortholog maps allow the findings to be placed in a broader biological context. The platform represents an innovation in genomics data visualization that will serve as a template for future tool development. Killiverse is freely accessible at https://killiverse.org/.

17
Minimally invasive measurement of maternal transcripts enables predicting the developmental potential of mammalian zygotes

Inoue, A.; Cheung, N.; Yamanouchi, T.; Matsuda, H.; Yoshioka, H.; Takeuchi, H.; Nishioka, M.; Yamamoto, M.; Wei, Y.; Houri, K.; Sato, H.; Guo, R.; Kamio, A.; Kobayashi, H.; Kono, T.; Matsumoto, K.; Miyamoto, K.

2026-06-06 developmental biology 10.64898/2026.06.02.726088 medRxiv
Top 0.6%
0.9%
Show abstract

Maternal transcripts are stored in the oocyte cytoplasm during oogenesis and play a pivotal role in early embryonic development after fertilization. However, specific maternal transcripts that reflect the developmental potential of embryos have not been systematically identified, and the use of maternal transcript levels as an indicator of successful development has not been explored. Here, we link the maternal transcriptome to the zygotes developmental potential by examining transcripts in a single polar body. The transcriptome of a zygote or an oocyte was highly similar to that of its accompanying polar body in mouse, cow, and human. We have identified a set of maternal transcripts whose expression levels fluctuate between poor- and good-quality zygotes. Specifically, Sipa1and Zmym6 were identified as marker transcripts that accurately reflect the developmental potential of zygotes. Using these marker genes, combined with machine learning, the development of zygotes to the blastocyst stage was successfully predicted with more than 80% specificity as early as 12 hours after fertilization. Furthermore, our prediction platform significantly improved implantation rates and live births to term. Thus, we have demonstrated a minimally invasive method for identifying maternal transcripts associated with zygote developmental potential. Our developed prediction system provides a generalizable conceptual framework for human infertility treatment to reduce the risk of implantation failure by excluding embryos with low developmental potential, especially when early embryos are transferred, and for livestock propagation to assess selected expressed maternal trait-associated variants before embryo transfer.

18
Radicular and periodontal structural defects underlie refractory oral pathology in the adult Hyp mouse model of X-linked hypophosphatemia

Nishizawa, C.; Miura, J.; Iwayama, T.; Yamazaki, M.; Michigami, T.; Miyagawa, K.

2026-04-30 pathology 10.64898/2026.04.27.719778 medRxiv
Top 0.6%
0.9%
Show abstract

ObjectiveX-linked Hypophosphatemia is associated with dental complications, including spontaneous endodontic infections (abscesses) in non-carious teeth and severe periodontal loss. Previous studies have mainly focused on dentin Hypomineralization; however, the structural basis underlying periodontal tissue failure remains unclear. We aimed to investigate histoanatomical abnormalities in the dentin and periodontium of Hyp mice to clarify structural consequences of Phex deficiency in adult molars. MethodsWe performed detailed histological and scanning electron microscopy analyses on the molar regions of untreated adult Hyp mice and wild-type littermates, with particular attention to the structural integrity of the root and periodontal ligament. Additionally, odontoblast process morphology and periodontal attachment abnormalities were evaluated. ResultsHyp molars exhibited marked root abnormalities, including radicular shunt-like defects and disorganized odontoblast processes, particularly in furcation and radicular dentin. Periodontal attachment showed characteristic asymmetry: detachment from the cementum surface was frequently observed, whereas attachment to the alveolar bone surface was relatively preserved. These changes were accompanied by thinning and discontinuity of Sharpeys fibers and increased vascularity in the periodontal ligament. ConclusionsThese findings provide a histoanatomical framework for understanding refractory dental complications in X-linked hypophosphatemia and support the importance of intervention during root development.

19
Increasing Phenomic Prediction Efficiency Using A Principal Component Analysis Based Pre-Processing Of Near Infrared Spectra

Bienvenu, C.; Roger, J.-M.; Sene, M.; Castro Pacheco, S. A.; Singer, M.; Felaniaina, B. L.; Terrier, N.; De Bellis, F.; Pot, D.; DE VERDAL, H.; Segura, V.

2026-05-13 genetics 10.64898/2026.05.10.724118 medRxiv
Top 0.7%
0.9%
Show abstract

Phenomic prediction (PP) is a breeding value prediction method using near infrared spectroscopy (NIRS). Spectra pre-processing is a key step in the analysis pipeline of PP and generally involves chemometrics methods. However, there is still little understanding in the genetics community of what pre-processing does and why it increases performances. Consequently, the choice of pre-processing is done either arbitrarily or through a search of the optimal set of methods and associated parameters. In this study, we propose a PCA-based pre-processing method where genetic values of spectra are estimated on a set of principal components instead of individual wavelengths. This way, estimations are based on a few informative and orthogonal features of spectra instead of many correlated, uninformative wavelengths. We tested this new pre-processing method on five data sets representing four plant species (maize, rice, sorghum and grapevine). Results show that it performs as good, or better than the best classical chemometric pre-processing methods in almost all cases. Combining PCA-based and classical chemometric pre-processing methods maximizes predictive ability. Moreover, this pre-processing method opens up possibilities of better understanding and selecting parts of the spectral information that are relevant for the prediction of breeding values. Indeed, components representing together about 1% of spectral variability were found to be responsible for most of PP predictive ability. Plain language summaryCultivated plants are the result of a breeding process during which their genetic values are used to select those to breed. Estimation of breeding values requires heavy experimental means and is time consuming. Phenomic prediction is a low cost and high throughput genetic value estimation method that is increasingly being used. It often uses near infrared spectroscopy measurements as predictors of genetic values that are easy to collect and thus routinely used in many species. However, near infrared spectra generally require pre-processing before being used in prediction. Currently used pre-processing methods arise from the chemometrics community, and still deserve a better in-depth appropriation by geneticists. In this study, we propose a new pre-processing approach that performs as good as or better than the best chemometric pre-processing generally used, reduces computation time, and allows for a better understanding of what parts of spectral information are relevant for prediction. Core IdeasO_LIWorking on principal components of spectra instead of wavelengths increases predictive ability of phenomic prediction and performs as good as or better than classical chemometrics pre-processing C_LIO_LIWorking on principal components of spectra requires less optimization of parameters than chemometrics pre-processing C_LIO_LIAbout 1% of spectral variance is responsible for most of the predictive power of phenomic prediction C_LIO_LIWorking on principal components of spectra pre-processed with classical chemometrics pre-processing can increase predictive ability even more C_LIO_LIPCA-based methods are valuable to optimize predictive ability of phenomic prediction and could be used more widely in the quantitative genetics field C_LI

20
Transcriptomic-guided compound prioritization and proteomics validation for HNRNPU deficiency identify signalling correction

Ye, X.; Tikhomirova, D.; Oksanen, M.; Gaetani, M.; Gharibi, H.; Mastropasqua, F.; Tammimies, K.

2026-05-07 molecular biology 10.64898/2026.05.04.722615 medRxiv
Top 0.7%
0.9%
Show abstract

Heterogeneous nuclear ribonucleoprotein U (HNRNPU) deficiency is a rare genetic cause of neurodevelopmental disorders (NDDs) lacking targeted therapies. Here, we developed a transcriptomic-guided compound prioritization pipeline using Connectivity Map (CMap) analysis on multi-model transcriptomic signatures from HNRNPU-deficient human cells and mouse models. Ten compounds were selected through manual curation and functionally screened in patient-derived HNRNPU-deficient neuroepithelial stem (NES) cells with earlier observed cellular phenotypes. Two of the compounds, AS601245 and Lenalidomide, significantly reduced the elevated neural progenitor population during differentiation, and their combination further decreased primary cilia incidence, indicating partial rescue of the patient-specific cellular phenotypes. To understand the mechanisms underlying the partial rescue, we employed proteome integral solubility alteration (PISA) and expression proteomics. PISA assay identified TMEM150C and GSK3A as proximal targets of combined treatment. Additionally, we observed reversal of multiple biological pathways including downregulation of Wnt signalling and upregulation of mitochondrial pathways and transmembrane proteins. Altogether, we established a computational-experimental pipeline for transcriptomic-guided drug repurposing for a monogenic NDD, and demonstrated that the network-level modulation partially rescues the delayed neural differentiation in HNRNPU-deficient neural cells.